The Journal of Molecular Diagnostics
○ Elsevier BV
Preprints posted in the last 30 days, ranked by how well they match The Journal of Molecular Diagnostics's content profile, based on 39 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Burssed, B.; van der Sanden, B.; Hops, W.; Neveling, K.; Kamping, E.; van Beek, R.; den Ouden, A.; Derks, R.; Timmermans, R.; Perrone, E.; Ramos, M. A.; Bellucco, F. T.; Hoischen, A.; Melaragno, M. I.
Show abstract
Complex rearrangements are one of the rarest types of structural variants (SVs) and can be divided into two categories: complex chromosomal rearrangements (CCRs) and complex genomic rearrangements (CGRs). CCRs include structural rearrangements that present at least three breakpoints and show exchange of genetic material between more than two chromosomes and CGRs are rearrangements that present more than one junction and/or more than one SV in cis. They are usually formed by one of the chromoanagenesis mechanisms, where a massive disruptive cellular event leads to multiple structural rearrangements. Classical cytogenomic techniques have been commonly applied for their characterization, but methodologies that involve longer DNA molecules, namely optical genome mapping (OGM) and long-read genome sequencing (lrGS), present a considerably higher SV detection resolution, revealing more details about the rearrangements, including precise breakpoint location. Here, we describe six patients with complex rearrangements investigated through a combination of different techniques: karyotyping, chromosomal microarray, and OGM were performed to characterize the rearrangements. Subsequently, lrGS was used to further resolve the alterations, refine their breakpoints' location, and sequence their junction points. Three patients presented CCRs involving three, four, and six chromosomes, while three exhibited CGRs involving one different chromosome each, providing a variety of complex SVs to show the importance of each technique and their combination in rearrangement resolution. In total, the complex rearrangements presented 127 breakpoints, 66 junction points and involved 14 of the 24 chromosomes. Higher-resolution techniques revealed additional complexity in all cases. Despite the advances provided by OGM and lrGS, conventional karyotyping remained indispensable for complete rearrangement resolution. In two patients, the findings supported a novel mechanism combining features of the different chromoanagenesis processes. Furthermore, evidence of inherited alterations was identified, and the comprehensive characterization of the rearrangements enabled more accurate genotype-phenotype correlations. Our findings indicate that an integrated approach combining karyotyping, OGM, and lrGS can completely resolve SVs, including complex rearrangements.
Yang, Y.; Vasudevaraja, V.; Serrano, J.; Mohamed, H.; Kelly, S.; Jour, G.; Gindin, T.; Park, K.; Jones, D.; Feng, X.; Pinnell, J.; Mclennan, S.; Tin, M. Y.; Tsirigos, A.; Snuderl, M.; Wrzeszczynski, K. O.
Show abstract
Next-generation sequencing (NGS) for the detection of somatic variants has become the method of choice in a variety of molecular oncology fields and in the clinic. Its use ranges from sequencing entire tumor genomes and transcriptomes to targeted clinical diagnostic gene panels. The NYU Langone Genome PACT (Profiling of Actionable Cancer Targets, LG-PACT) assay is a qualitative in vitro diagnostic test that uses targeted next generation sequencing (NGS) of formalin-fixed paraffin-embedded (FFPE) tumor tissue matched with normal specimens from patients to detect gene alterations in a targeted panel covering 606 genes and the TERT promoter. Indications for testing are cancer (solid tumors and hematological malignancies) where a mutational profile from multiple genes would be informative for disease stratification, prognosis, or treatment options including targeted therapies and eligibility for clinical trials. The test is intended to provide information on somatic mutations including point mutations, small insertions/deletions (indels), and copy number aberrations for diagnostic and treatment decisions. LG-PACT is a United States Food and Drug Administration (FDA) cleared diagnostic test (510K: K202304). The clinical interpretation of sequencing data of molecular tumor markers from NGS encompasses automated variant calling tools with human interpretation. This final mostly manual review of data step is intensive, involving highly trained scientists, encompassing literature review, interpretation and clinical tier classification by pathologists, who then provide a complete molecular diagnostic report to the treating oncologists. We provide analysis of 1339 clinical genomic profiles from 31 different cancers and their subtypes, comprising of central nervous system (CNS) 792 (59%) cases (incl. meningioma, glioma and glioblastoma), with 267 (20%) cases predominantly of lung, pancreatic and colorectal and 280 of others (21%). Here, we present the technical challenges of validating an NGS oncological diagnostic targeted assay for clinical grade accuracy and sensitivity for patient care. We show how copy number alterations provide a more comprehensive description of the tumors genomic profile. We then outline the utility of targeted panel sequencing based on certified pathologist selection of reportable variants for our current patient cohort. Where analysis of variant detection has led to 49.4% (661/1339) of our clinical tumor samples containing mutations in known therapy targeted genes, 35.6% (477/1339) with mutation detected in other genes, and 15% (201/1339) cases being negative.
Ayati, A.; Onal, G.; Sur, A.; Azzam, S.; Wang, B.; Rudrapatna, V. A.
Show abstract
Objective: Erythropoietic protoporphyria (EPP) is a rare photodermatosis marked by multi-year diagnostic delays. We developed and externally validated machine learning models to identify patients with EPP earlier from longitudinal electronic health record (EHR) data and estimate undiagnosed disease burden. Materials and Methods: In a retrospective case-control study at two San Francisco health systems, an academic referral center (UCSF) and a safety-net hospital (ZSFG) we identified 74 confirmed EPP cases using combined diagnostic coding, biochemical criteria, and specialty chart review. Symptom-enriched controls were sampled at a 40:1 ratio. Longitudinal diagnoses, laboratory results, medications, procedures, and encounters preceding the outcome date were modeled with a gradient-boosting classifier (CatBoost) and a state-space sequence model (MAMBA). The best model was deployed across the UCSF population and externally validated at ZSFG without retraining. Results: On the UCSF held-out test set (n=1,865; 43 cases), MAMBA outperformed CatBoost (AUC ROC 0.91 vs 0.89; average precision 0.42 vs 0.27; precision 65% vs 20%), flagging cases a median of 229 days before documented diagnosis. Deployed across 297,967 symptom-compatible patients, it identified 310 high-risk individuals, implying a prevalence approaching genetic estimates. External validation at ZSFG showed attenuated performance (AUC ROC 0.72; average precision 0.10) while preserving early detection (median 264 days). Discussion: A sequence model integrating temporal EHR signals detected EPP months before clinical recognition, corroborating genetic evidence of substantial underdiagnosis. Cross-site attenuation reflects population and documentation differences and underscores the need for local recalibration. Conclusion: Longitudinal EHR-based machine learning can shorten EPP diagnostic delay and prioritize patients for confirmatory testing, supporting proactive rare-disease case finding.
Lane, T.; Green, T. E.; Garza, D.; Brown, N. J.; de Silva, M. G.; Bennett, M. F.; Tubb, C.; Macdonald, S. M. W.; Gascoigne, A.; Phillips, R. J.; Slavin, J.; D'Arcy, C.; MacGregor, D.; Clifford, A.; Pathmanathan, L.; Robertson, S. J.; Bekhor, P.; Simpson, J.; Gooley, S.; Scheffer, I. E.; Berkovic, S. F.; Penington, A. J.; Hildebrand, M.
Show abstract
Targeted precision therapies are increasingly used in the treatment of individuals with vascular anomalies (VAs). This increases the need for rapid, accurate and inexpensive genetic diagnosis. Droplet digital polymerase chain reaction (ddPCR) is an alternative to next-generation sequencing (NGS), permitting rapid, highly sensitive interrogation of recurrent pathogenic mosaic variants. We examined the feasibility of ddPCR as a primary diagnostic tool in a large cohort of individuals with VAs. Lesional tissue was collected for ddPCR of up to 46 recurrent pathogenic variants across 16 genes associated with VAs. Specimens were assessed on a subset of assays for each individual based on clinical phenotype. Most individuals who had negative ddPCR results went on to high-depth gene panel or deep exome NGS, or Sanger sequencing. Here we report the phenotypic and molecular findings for 78 newly recruited and tested individuals in addition to the 60 individuals already reported from our cohort. The overall diagnostic yield for our cohort when combined with individuals previously reported was 104/138 (75%). Of 138 individuals tested, recurrent pathogenic variants were detected in 71 (51%) on ddPCR. Variants were most frequently identified in PIK3CA (n=28), TEK (n=18), GNAQ (n=12), or MAP2K1 (n=7). In a further 33 individuals, pathogenic variants were identified on NGS or Sanger sequencing. Our findings indicate that ddPCR is an efficient method achieving a high diagnostic yield in our cohort when used prior to sequencing.
Cornelli, L.; Nhat Nguyen, T.; Van Belle, R.; Roelandt, S.; De Cock, A.; Van der Meulen, J.; Loontiens, S.; Van Roy, N.; De Preter, K.
Show abstract
An important step toward clinical implementation of (epi-)genomic assays on liquid biopsies is their validation on identical samples within and across laboratories. For these validation studies, there is a need for cell-free DNA (cfDNA) samples with defined tumor fractions and (epi-)genomic aberrations. However, the amount of circulating cfDNA isolated from patient samples is often limited, especially in pediatric cases. Additionally, patient samples contain a high degree of variability in cfDNA yield and tumor fraction. Several commercial artificial cfDNA products are available for validation studies, however their use is restricted to specific assays, aberrations and/or tumor entities. Alternatively, artificial cfDNA samples can be produced by fragmenting genomic DNA to mimic highly fragmented cfDNA derived from both tumor and healthy blood, followed by mixing artificial tumoral and healthy cfDNA at defined fractions. In this study, we compared native cfDNA with artificial cfDNA generated by three different fragmentation methods, including sonication and two enzymatic digestions using micrococcal nuclease and double-stranded deoxyribonuclease (dsDNase). We assessed fragment length profiles, end motifs and nucleosome occupancy patterns from shallow whole-genome sequencing data, as well as coverage profiles from targeted panel sequencing, together with a small-scale mixing experiment of tumor and healthy cell derived artificial cfDNA. Although sonication remains a convenient high-throughput approach to generate artificial cfDNA for certain downstream applications, enzymatic fragmentation, particularly the dsDNase-based method, more faithfully reproduced native cfDNA characteristics.
Viz-Lasheras, S.; Dacosta, A.; Rivero-Calle, I.; Martinon-Torres, F.; EUCLIDS, GENDRES, PERFORM, and DIAMONDS consortia, ; Gomez-Carballa, A.; Salas, A.
Show abstract
Accurate discrimination between viral, bacterial, and inflammatory diseases in febrile children remains a major clinical challenge that contributes to diagnostic uncertainty, inappropriate antimicrobial use, and suboptimal clinical management. Host blood transcriptomics offer a promising strategy to improve diagnostic precision. The present study represents the largest integrative multi-cohort pediatric study of transcriptomic biomarker discovery, validation, and confirmation reported to date, integrating harmonized public transcriptomic datasets with an independent confirmation cohort comprising well-phenotyped patients to identify parsimonious host-response signatures for differentiating viral, bacterial, and inflammatory diseases. Transcriptomic signatures were derived from an integrated retrospective microarray multi-cohort (n=1,683), independently validated in a retrospective RNA-seq cohort (n=767), and confirmed by digital PCR in an independent cohort (n=29), demonstrating reproducibility across patient populations, transcriptomic technologies, and analytical platforms. The analysis identified binary signatures and a unified multiclass classifier that consistently achieved high diagnostic accuracy across all three study phases and outperformed more than 30 published host transcriptomic signatures. Decision curve analysis showed substantially greater clinical net benefit than C-reactive protein across clinically relevant decision thresholds. These findings provide a strong foundation for clinically deployable molecular diagnostics to improve patient triage, antimicrobial stewardship, and precision medicine in childhood infections.
Guedes, J.; Sliwa-Gonzalez, A.; Szadai, L.; Geiger, P.; Woldmar, N.; Reyes, M. A.; Bastida, R. A.; Coto, D. L. F.; Oskolas, H.; Marko-Varga, M.; Schultz, L.; Appelqvist, R.; Wieslander, E.; Malm, J.; Marko-Varga, G.; Gil, J.
Show abstract
Melanoma incidence continues to rise globally, with formalin-fixed paraffin-embedded (FFPE) tissue archives representing an invaluable resource for large-scale retrospective proteomic studies. However, inconsistent deparaffinization remains a critical pre-analytical bottleneck limiting protein yield, reproducibility, and downstream data quality. In this study, we developed and validated a fully automated FFPE deparaffinization workflow using the Fluent(R) 780 liquid handling workstation (Tecan (C)) and evaluated its performance against a conventional manual protocol in a cohort of 54 patients with primary cutaneous melanoma, predominantly at early AJCC 8th edition stage I-II. The automated workflow achieved superior protein identification (6,146 {+/-} 860 vs. 4,941 {+/-} 1,091 proteins; p < 0.0001) with lower technical variability, while maintaining highly comparable global proteomic profiles as confirmed by principal component analysis and hierarchical clustering. A total of 8,305 proteins (96.1%) were identified by both methods, supporting the reproducibility and equivalence of the automated approach. Patients were stratified by the presence (N=21) or absence (N=33) of histological regression in the primary tumor. Proteomic comparison revealed 97 upregulated and 226 downregulated proteins in regressing melanomas, with pathway enrichment analysis demonstrating elevated mitochondrial and translational activity alongside reduced innate immune and complement pathway activation in the regression group. No statistically significant differences in overall, disease-free, or progression-free survival were observed between groups, consistent with the early-stage composition of the cohort. Digital pathology validated tissue morphology preservation across processing conditions. These findings support the integration of automated FFPE processing with proteomic and digital pathology workflows as a scalable platform for precision melanoma research. TOC Figure O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=133 SRC="FIGDIR/small/744404v1_ufig1.gif" ALT="Figure 1"> View larger version (49K): org.highwire.dtl.DTLVardef@1d51629org.highwire.dtl.DTLVardef@a1f126org.highwire.dtl.DTLVardef@1df1b0aorg.highwire.dtl.DTLVardef@686f1c_HPS_FORMAT_FIGEXP M_FIG C_FIG
Płonka, W.; Kostka, D.; Lalik, A.; Kurpas, M.; Dinh, K. N.; Sitkiewicz, M.; Kimmel, M.; Rzyman, W.; Jaksik, R.
Show abstract
Formalin-fixed, paraffin-embedded (FFPE) tissues remain an essential resource for molecular studies, yet formalin-induced cytosine deamination introduces characteristic C>T/G>A artifacts that compromise the accuracy of next-generation sequencing (NGS) analyses. Numerous computational methods and enzymatic DNA repair strategies have been proposed to reduce these artifacts, but no systematic comparison across tools and experimental conditions exists. Here, we evaluate the performance of seven computational approaches (SOBDetector, Ideafix, MicroSEC, FFPolish, DeepOmics FFPE/FFPE-PLUS, FFPErase) together with the NEBNext(R) FFPE DNA Repair Mix v2, a multi-enzyme repair system applied during DNA preparation. Using three independent datasets, one based on whole genome sequencing (CGCI-BL) and two on whole exome sequencing (TCGA-PC and SUT-LUAD, the latter containing enzymatically repaired samples), and matched fresh-frozen samples as the gold standard, we assess precision, sensitivity, and artifact reduction efficiency across all methods. We further examine the potential synergy between enzymatic repair and post-sequencing computational filtering. Our results provide practical guidelines for FFPE artifact correction and demonstrate that enzymatic treatment provides the best results, while among the computational methods, FFPErase offers the most robust reduction of cytosine deamination artifacts while maximizing the retention of true somatic variants. KEY MESSAGESO_LIFormalin fixation in FFPE samples introduces artifacts that can significantly affect the accuracy of NGS analyses. C_LIO_LIAmong the evaluated approaches, enzymatic repair using NEBNext(R) FFPE DNA Repair Mix v2 achieves the most effective reduction of sequencing artifacts. C_LIO_LIComputational methods vary in performance, with FFPErase showing the most robust balance between artifact removal and retention of true somatic variants. C_LIO_LICombining enzymatic repair with computational filtering did not lead to consistent improvements in performance across datasets. C_LI
McMahon, K.; Nielsen, S.; Knoll, H.; Talwar, R.; Thompson, D.; Wilkason, C.; Ozonoff, A.; Stachler, E.; Sabeti, P.
Show abstract
The 2026 Bundibugyo ebolavirus (BDBV) outbreak underscores the need for rapidly deployable molecular diagnostics. We developed and analytically validated reverse-transcription quantitative PCR assays detecting BDBV, Zaire ebolavirus, and Sudan ebolavirus. The platform includes a BDBV singleplex assay, a duplex assay with a human internal control, a four-target multiplex assay for ebolavirus differentiation, and a probe-free SYBR Green assay. We adapted the assays to a portable qPCR instrument, reducing runtime from 65 to 35 minutes, and validated lyophilized reagents to reduce cold-chain requirements. All TaqMan formats achieved a 95% limit of detection of 5 copies per reaction across instruments and reagent types; the SYBR Green assay achieved 50 copies per reaction. The assays detected viral RNA in contrived clinical samples without cross-reactivity among ebolavirus species tested. We shared the protocols in real time through Ampliphi (https://www.ampliphi.bio), a new open-access platform for rapidly disseminating diagnostic assays, and through protocol.io.
Zhu, M.; Li, A.; Safa, I.; Galera, P.; Hazoglou, M.; Vanderbilt, C.; Kamali, A.; Goldgof, G.; Veeraraghavan, H.; Jiang, J.; Ardon, O.; Geneslaw, L.; Dogan, A.
Show abstract
Pathologic diagnoses of hematopoietic diseases require immunohistochemistry (IHC) stains selected by pathologists upon preview of H&E-stained slides. This multi-step workflow can delay diagnostic turnaround time by days. Hence, we developed the Hematopathology Automatic Triaging System (HATS), which automates IHC panel ordering directly from H&E whole-slide images using pretrained pathology foundation model representations combined with attention-based multiple-instance learning. After the most comprehensive evaluation of pathology foundation models for hematologic malignancy classification to date, encompassing seven publicly available models, we trained HATS on 4,996 whole-slide images from 1,607 patients spanning the ten most common lymphoma diagnostic categories. HATS achieves 84% case-level subtype classification accuracy (0.962 ROC-AUC), translating to 92% IHC panel ordering accuracy. In a blinded reader study, HATS outperforms practicing pathologists at predicting lymphoma subtypes from morphology alone (85% vs 65%). In an independent real-world validation of 230 clinical cases, after directing 7 cases with scant tissue for manual review, HATS-ordered IHC panels were sufficient for diagnosis in 72.6% of cases. By automating the triaging step while preserving full pathologist oversight, HATS offers a safe and practical entry point for clinical AI adoption in pathology.
Yelmen, B.; Hofmeister, R. J.; Lutsar, V. K.; Finianos, M.; Stone, B. C.; Joeloo, M.; Krebs, K.; Kivistik, P. A.; Smit, S.; Estonian Biobank Research Team, ; Metspalu, M.; Hudjashov, G.; Milani, L.
Show abstract
Since copy number variations (CNVs) in pharmacogenes can cause significant alterations in drug metabolism, their reliable detection is of high importance both for large-scale studies and personalized medicine. Whole-genome sequencing, and specifically long-read sequencing, is the gold standard for CNV detection. Despite increasing availability of these technologies, genotyping arrays are still widely used as cost-effective alternatives in biobank and clinical settings, yet calling CNVs based on array intensity signals is challenging due to low base pair resolution. In this work, we developed a neural network model, nnCNV, to predict deletions in the CYP2C19 pharmacogene region from array intensity signals. We compared our method to the most widely used algorithm, PennCNV, and demonstrated better performance reaching 100% accuracy in the test dataset. Furthermore, we predicted probe-by-probe CYP2C19 deletion coordinates for all Estonian Biobank samples using nnCNV and PennCNV, and validated these predictions using an identity-by-descent (IBD) sharing method, which also demonstrated superior nnCNV performance. For the deletion samples with conflicting PennCNV and nnCNV predictions, we performed PCR analysis for validation, which showed 97% precision for nnCNV compared to 23% for PennCNV. Finally, we assessed the gradient-based feature importance maps and showed that nnCNV utilizes signal intensity information not only from deletion probes, but also from probes in flanking regions. Our results demonstrate that long-range information, which cannot be utilized by hidden Markov models, can improve CNV calling.
Meerson, A.
Show abstract
To explore adapting qPCR systems for end-point nucleic acid quantification using dyes such as SYTO-9, we quantified serial dilutions of DNA and RNA standards in the range of 0.75 - 200 ng/{micro}l on 384-well qPCR devices. SYTO-9 fluorescence was successfully measured using standard SYBR Green settings. Blank-subtracted relative SYTO-9 signal showed a logarithmic dependence on DNA/RNA concentration (R2 > 0.95). Measurements were highly stable with different incubation times, temperatures of up to 95{degrees}C, and photobleaching. The described approach is a valuable QC option for high-throughput DNA/RNA isolations and could be adapted to additional fluorometric assays beyond nucleic acids.
Pradani, G. A. P.; Alifia, A.; Syahbaniati, A. P.; Larasmanah, A. N.; Busaeri, M.; Djunaedy, H.; Choerunisa, T. F.; Massi, M. N.; Rachman, R. W.; Fibriani, A.; van Crevel, R.; van Ingen, J.; Lestari, B. W.
Show abstract
As drug-resistant tuberculosis (DR-TB) cases rise, resistance detection in a timely manner is essential to lead effective treatment and limit transmission. Targeted next-generation sequencing (tNGS) offers quick results with multiple important drugs covered, but assessments regarding its performance for DR-TB diagnostic use compared to whole genome sequencing (WGS) as the most comprehensive genomic-based tool are still limited. This cross-sectional study compared resistance profiles generated by Deeplex Myc-TB tNGS assay with WGS for 116 prospectively-collected rifampicin resistant TB samples from West Java, Indonesia. All 116 samples were subject to paired analysis, the clinical samples were split to be directly processed for tNGS and to be cultivated for culture-based WGS. Both WGS and tNGS were carried out using Illumina MiSeq platform. High concordance of tNGS and WGS were observed across thirteen anti-TB drugs evaluated, particularly for drugs included in the BPaLM regimen. Isoniazid had the lowest concordance of 86.73%. Of 116 samples, 31.03% (n = 36) had discrepant resistance calling from the two methods for one or more drugs, which came from 73 discordant variants identification. The most common source of discrepancy was when tNGS detected a resistance-conferring mutation while WGS did not (54.8%). tNGS could detect mixed infection better than WGS, but WGS was superior in identifying detailed major Mycobacterium tuberculosis lineage of the sample. tNGS showed a good level concordance with WGS in detecting resistance-conferring mutations in rifampicin-resistant TB samples, with a more rapid turnaround time. Continuous update to tNGS panel and mutation catalogue is needed to keep the tool clinically relevant. ImportanceDrug-resistant tuberculosis (DR-TB) continues to pose worldwide threat, and newer diagnostic tools to generate quick, comprehensive resistance profile are crucial to provide timely appropriate treatment. Targeted next-generation sequencing (tNGS) is a promising new alternative, but more evidence on its performance is needed to support programmatic adoption. By analysing DR-TB samples with both tNGS and whole genome sequencing (WGS) and evaluating their results agreement, this study shows that tNGS works just as well as WGS in detecting TB drug resistance-conferring mutations, confirming its potential for routine diagnostic use. This study also observed that while WGS is superior in identifying Mycobacterium tuberculosis lineage with high resolution, it did not detect mixed infection better than tNGS. Notably, this study demonstrated that tNGS is clinically relevant for DR-TB detection in a high burden setting, providing evidence for programmatic consideration in Indonesia and other settings with similar demographics and TB situation.
Lee, K. T.; Egleston, B.; Fetzer, D.; Domchek, S. M.; Fleisher, L.; Wen, K.-Y.; Wagner, L.; Roberts, S.; Howe, S.; Cacioppo, C.; Christiansen, J.; Karpink, K.; Selmani, E.; Mastaglio, E.; Weinberg, M.; Wood, E. M.; Feng, J.; John, S.; Schweickert, K.; Mcleod, B.; Bradbury, A. R.
Show abstract
Background: Many at-risk patients lack access to genetic services due to a genetic counselor (GC) workforce shortage. Little is known about how digital alternatives impact patients with and without cancer who meet criteria for genetic testing. Methods: eREACH2 is a randomized 4-arm non-inferiority trial where pre-test (visit 1) and/or return of results (visit 2) GC counseling was replaced with a patient-centered digital intervention. Arms include: A (GC/GC), B (GC/digital), C (digital/GC) and D (digital/digital). Primary outcomes were non-inferiority in uptake of genetic services and change in genetic knowledge and general anxiety from baseline to post-disclosure of results (T0-T2). Secondary cognitive and affective outcomes were assessed using non-inferiority ANOVAs and equivalency chi-squared tests in intention-to-treat and per-protocol analyses. Findings: 773 participants were recruited nationwide; 46.6% from rural areas. Mean age was 51 years (range 20-87), 13% male, 12% non-white, 29% had less than a college education, and 33% had a personal history of cancer. 584 (76%) patients completed testing (14% had a positive result, 16% had a VUS). In the primary ITT analyses, we met the non-inferiority for uptake of genetic services and anxiety, but results were inconclusive for knowledge. Secondary outcomes were heterogeneous across arms. Arm C demonstrated consistently favorable effects, while Arms B and D showed less favorable outcomes in select domains (e.g. satisfaction and MICRA). Patients who received positive or VUS results via digital disclosure had significantly higher MICRA scores - indicating greater negative response to testing. Interpretation: In this large, randomized trial of patients with and without cancer, the eREACH intervention was effective for pre-test counseling, but inconclusive for digital disclosure of results. Exploratory analyses suggest that digital delivery could be a reasonable alternative for individuals receiving negative results, while those receiving positive or VUS results may derive some short-term psychosocial benefit from GC disclosure.
Gordon, D. C.; Thumbadoo, K. M.; Naidoo, S.; Nishimura, A. L.; Rodrigues, M.; Fraser, H.; Cutrupi, A. N.; Roxburgh, R. H.; Shaw, C. E.; Kennerson, M. L.; Scotter, E. L.
Show abstract
Pathogenic missense variants in the X chromosome gene UBQLN2 cause amyotrophic lateral sclerosis (ALS), often accompanied by frontotemporal dementia (FTD). As an X-linked gene, UBQLN2 is subject to X chromosome inactivation (XCI), a process wherein one X chromosome in each cell is randomly inactivated to a Barr body throughout the body in females, creating a mosaic of allelic expression in the tissues of heterozygotes. Despite heterozygous females constituting a majority of reported cases of UBQLN2-linked ALS/FTD, and the known influence of XCI on neurological disorders at large, no current disease models account for XCI. Here we report the characterisation of 12 iPSC clones carrying the ALS/FTD-causing p.T487I (c.1460C>T) UBQLN2 variant. These clones, originally derived from 3 heterozygous carrier fibroblast lines, underwent validation of homeostatic Barr body retention. Erosion of XCI in a subset of the lines was correlated with biallelic expression (of both wildtype and mutant UBQLN2), as measured through a novel allele-selective qPCR (AS-qPCR) assay and verified by amplicon-based Illumina sequencing and Sanger chromatogram quantification, enabling selection of iPSC clones best retaining XCI. Together, this UBQLN2 AS-qPCR assay and selected iPSC clones will enable studies of the role of XCI and its skew in female resilience to UBQLN2 p.T487I-linked ALS/FTD and enable development of allele-selective therapies.
Arif, A.; Filho, J. V. d. S.
Show abstract
The increasing use of tumor sequencing has intensified the need for fast, traceable interpretation of genomic variants. General-purpose large language models can produce fluent answers, but unsupported statements, weak provenance, and stale knowledge limit their suitability for clinical genomics. We developed OncoGenRAG, a research framework that combines a parameter-efficiently fine-tuned BioBERT classifier with an entity-aware retrieval system over a curated, multi-source oncology knowledge base. The reported knowledge base contains 933 harmonized records derived from CIViC, ClinVar/dbSNP, Open Targets, UniProtKB/Swiss-Prot, Ensembl Variation, and linked PubMed literature. The classifier assigns one of five labels: Pathogenic, Likely Pathogenic, Variant of Uncertain Significance, Benign, or Oncogenic; the retrieval component ranks evidence records using subword TF-IDF similarity and explicit gene, variant, and cancer-type matches. A rejection rule suppresses answers when retrieval support is below a prespecified threshold. In the authors held-out evaluation, the classifier achieved 92.40% accuracy, 93.15% weighted precision, 92.40% weighted recall, and 92.65% weighted F1 score. In a separate benchmark of 100 clinical-style queries, OncoGenRAG achieved reported Precision@1 of 94.5%, Precision@3 of 96.8%, and 100% database grounding. No hallucinated answer was observed under the study operational definition, compared with a 41.0% no-hallucination rate for the ungrounded baseline. These results should be interpreted as internal validation rather than proof of universal safety because query construction, annotator agreement, class-specific performance, calibration, and external validation data were not available for independent analysis. OncoGenRAG provides a transparent design for evidence retrieval and abstention, but it is a research prototype and must not be used to select treatment without expert review.
Hodel, F.; Thorball, C. W.; Haefliger, D.; Cerutti, L.; Cattaneo, P.; Howald, C.; Männik, K.; de La Harpe, R.; Samer, C. F.; Xenarios, I.; Fellay, J.; Girardin, F. R.
Show abstract
Background. Pharmacogenetic (PGx) testing can guide drug prescribing but remains limited by the genomic assay used. Genotyping arrays are widely implemented yet limited to predefined variants, whereas low-pass whole-genome sequencing (LP-WGS) is not constrained by fixed probe design and may provide broader PGx variant availability after imputation. Methods. We compared Illumina Global Screening Array (GSA) v3 with ~1x LP-WGS for PGx profiling in 500 hospital biobank participants with electronic health record evidence of exposure to pharmacogenetically actionable drugs and reported adverse drug reactions. Concordance was evaluated genome-wide, at 20 actionable pharmacogenes for PharmCAT-derived star alleles and metabolizer phenotypes, and for HLA alleles. Results. Genome-wide concordance between imputed array and LP-WGS data was high (median 99.63%; interquartile range, 99.59%-99.64%). For pharmacogenetically relevant variants, LP-WGS captured a larger fraction, particularly rare alleles absent from the array data, whilst maintaining high concordance at shared sites. Predicted phenotype concordance exceeded 98% for most genes, although gene-specific differences in phenotype classification were observed. LP-WGS reduced missing phenotype assignments for selected loci, particularly CYP2C19 and NAT2, by improving resolution of star-allele structure. However, in structurally complex or incompletely characterized genes such as CYP2C9 and CYP2D6, broader variant recovery increased indeterminate classifications rather than consistently improving clinical interpretability. For HLA loci, concordance varied by imputation strategy, with SNP2HLA performing marginally better utilizing the GSA array compared to the LP-WGS approach. Conclusions. Overall, LP-WGS provides broader variant coverage and improved resolution for selected pharmacogenes but did not resolve all clinically important loci. These findings support further evaluation of LP-WGS as a scalable PGx screening approach, especially where long-term genomic data reuse is a priority.
Buianova, A. A.; Cheranev, V. V.; Kuznetsov, M. I.; Repinskaia, Z. A.; Belova, V. A.
Show abstract
Introduction: The application of pharmacogenomics (PGx) in pediatrics is limited by the lack of age-oriented interpretation approaches, as algorithms developed for adults do not account for ontogenetic changes in the activity of drug-metabolizing enzymes and transport proteins. The aim of this study was to evaluate the clinical applicability of pharmacogenomic data in Russian children, assess the concordance between genotype-based recommendations and the ontogenetic status of drug-metabolizing enzymes, and develop recommendations for the generation of age-oriented PGx reports. Methods: We analyzed whole-exome sequencing (WES) data from 524 pediatric patients and 635 newborns, filtering pharmacogenomic annotations according to PharmGKB/ClinPGx evidence levels (1A-2B) and the presence of the 'Pediatrics' tag. The concordance between genotype-based recommendations and the ontogenetic status of drug-metabolizing enzymes was assessed in newborns. In a pediatric subgroup of 100 patients, a retrospective analysis of medical records was performed to evaluate the structure of pharmacotherapy and the frequency of adverse drug reactions (ADRs). A 'PGx-ADR-cost' database was created, and the relative population burden index was calculated for 27 gene-variant-drug-ADR associations. Results: Clinically relevant annotations (requiring drug avoidance or dose modification) accounted for only 5% of all initial pharmacogenomic annotations in both cohorts; 67.6% (pediatric cohort) and 67.2% (neonatal cohort) of these were related to alleles with altered function. Concordance between genotype-based recommendations and the ontogenetic status of drug-metabolizing enzymes in newborns was observed in only 5 of 14 (35.71%) gene-drug pairs. ADRs were identified in 21% of the 100 pediatric patients; however, only two cases could be explained by high-evidence PharmGKB/ClinPGx annotations. Ranking by relative population burden identified UGT1A1*28-irinotecan-induced neutropenia and HLA-A*31:01-carbamazepine-induced severe cutaneous reactions as priority associations. Conclusions: Age represents a critical factor in the interpretation of pharmacogenomic data in children, as current approaches to PGx reporting do not adequately incorporate the ontogenetic context. We propose a pediatric PGx interpretation model that includes mandatory reporting of patient age, ontogenetic adjustment, evidence-level stratification, and multidisciplinary clinical assessment. Prospective validation is required to confirm the clinical utility of the proposed approach.
Dashti, N.; Schneider, M. M. K.; Eckardt, J. N.; Fiebig, F.; Schweigler, D.; Buttner, S.; Middeke, J. M.; Bornhauser, M.; Rollig, C.; Kather, J. N.; Wiest, I. C.
Show abstract
Background: Adverse event (AE) coding is essential for safety monitoring in oncology clinical trials, particularly in acute myeloid leukemia (AML), where intensive therapies are associated with frequent and heterogeneous toxicities requiring standardized MedDRA (Medical Dictionary for Regulatory Activities) coding. However, manual Low-Level Term (LLT) assignment remains labor-intensive, subjective, and difficult to scale. Although large language models (LLMs) have emerged as promising decision-support tools for automated coding, unguided zero-shot generation remains insufficient for reliable fine-grained MedDRA coding. Objective: To develop and evaluate a retrieval-augmented reasoning pipeline for clinically aligned LLT-level MedDRA coding of free-text adverse events from prospective AML clinical trials. Methods: We implemented a retrieval-augmented reasoning pipeline inspired by the retrieval-augmented generation (RAG) paradigm using LLaMA-3.3-70B-Instruct as the primary backbone and benchmarked the framework across multiple open instruction-tuned LLMs. Dense semantic retrieval first generated a constrained top-100 LLT candidate set for each AE, followed by structured LLM reasoning to select a single best-matching LLT and deterministic mapping to Preferred Term (PT) and System Organ Class (SOC) levels. The pipeline was evaluated retrospectively on AE datasets from three prospective AML clinical trials (MOSAIC, DELTA, and DaunoDouble) with automated LLT/PT/SOC metrics and expert-assessed Clinical Correctness Rate (CCR). Results: Clinical expert review showed high clinical acceptability of the RAG pipeline across datasets (91-97%). Under automated evaluation, the pipeline achieved LLT exact accuracy of 50-58%, PT accuracy of 78-85%, and SOC accuracy of 90-93%. Zero-shot generation and random candidate selection performed substantially worse. Semantic retrieval more often included the coder-assigned LLT among the candidate terms available to the model than retrieval based on lexical similarity. Multi-model benchmarking showed that backbone choice mainly affected LLT exact agreement, whereas PT and SOC performance remained comparatively stable. Conclusions: Retrieval-augmented reasoning supports clinically aligned MedDRA coding of free-text adverse events under realistic candidate constraints in AML clinical trials. Evaluation across three AML clinical trials showed that strict LLT-level string agreement underestimated clinical ap-propriateness, highlighting the importance of combining hierarchical evaluation metrics with clini-cal expert validation for AI-assisted MedDRA coding in hematology trials.
De Keyzer, L.; Deserranno, K.; Skevin, S.; Van Hoofstat, D.; Deforce, D.; Van Nieuwerburgh, F.
Show abstract
Recombinase polymerase amplification (RPA) enables rapid nucleic acid testing in low-resource environments, but poorly characterized byproducts can compromise assay specificity and cause false-positive results. Here, we amplified the thirteen original CODIS core loci and Amelogenin to characterize recurrent RPA artefacts and establish conditions that reduce their formation. First, RPA products were analyzed for two reference samples by Oxford Nanopore Technologies sequencing. This revealed two distinct classes of multimeric products: primer multimers and amplicon multimers, consisting of repeated primer or amplicon sequences, respectively. Individual artefacts contained up to 281 primer copies or 22 amplicon copies, demonstrating the extensive range of these products. Next, we performed an optimization study to evaluate the effects of reaction temperature and reagent concentrations at two representative loci, D3S1358 and D5S818. Among the conditions tested, temperature had the most pronounced effect. Reducing the temperature from 42{degrees}C to 34{degrees}C increased the relative target amplicon fraction from 15% to 83% for D3S1358 and from 84% to 98% for D5S818, while maintaining or increasing absolute target concentration. Lower primer concentrations and higher T4 UvsX concentrations also reduced multimer formation, although lower primer concentrations reduced target yield and caused allelic dropout. Finally, amplification at 34{degrees}C was evaluated across all fourteen loci by sequencing. Relative to 42{degrees}C, the target read fraction increased by more than 5 percentage points for 7/14 loci in one reference sample and 9/14 loci in the other, with the largest improvements at multimer-prone loci. These findings identify multimers as an important class of RPA artefacts and establish reaction temperature and T4 UvsX concentration as promising conditions to improve RPA specificity.